今天會繼續使用 Books to Scrape 作為練習網站,實作以下內容:
使用 try...except 處理不同類型的錯誤。
使用 raise_for_status() 檢查 HTTP 請求是否成功。
使用 dict.get() 和條件判斷處理缺少的欄位。
驗證價格、評分和書名是否符合預期。
記錄錯誤資料,避免單筆錯誤影響整個爬蟲。
將成功資料與錯誤紀錄分開儲存。
一、認識爬蟲常見的錯誤
在實際爬取網站時,錯誤不一定只來自程式本身,也可能來自網路、伺服器或網頁內容。
以下是今天會處理的幾種情況:
錯誤類型
可能發生的原因
連線錯誤
網路不穩定或網站暫時無法連線
HTTP 錯誤
伺服器回傳 404、500 等錯誤狀態碼
欄位缺失
網頁沒有預期的 HTML 元素
資料格式錯誤
價格不是有效數字,或評分格式不符合預期
資料內容異常
書名為空、價格小於零等
這些問題可能導致程式停止,也可能讓錯誤資料混進最後的資料集。
所以今天會把爬蟲的流程拆成兩個部分:
第一部分是資料擷取: 確認網頁可以正常取得,並解析需要的欄位。
第二部分是資料驗證: 確認取得的資料符合預期格式,再決定是否要儲存。
二、實作一:使用 try...except 處理 HTTP 請求錯誤
前幾天已經使用過 try...except,今天會進一步把它應用在爬蟲的 HTTP 請求中。
先從最基本的請求開始。
import requests
url = "https://books.toscrape.com/"
try:
response = requests.get(url, timeout=10)
response.raise_for_status()
print("請求成功")
print("狀態碼:", response.status_code)
except requests.exceptions.RequestException as e:
print("請求失敗:", e)
這段程式碼中有幾個重要的部分。
requests.get() 負責向網站發送 GET 請求,而 timeout=10 表示設定請求逾時時間,避免程式無限期等待回應。
response.raise_for_status() 則會檢查 HTTP 狀態碼。如果伺服器回傳 4xx 或 5xx 狀態碼,就會產生 HTTP 錯誤。
最後,except requests.exceptions.RequestException 可以捕捉 Requests 常見的請求錯誤,例如連線失敗、逾時或 HTTP 錯誤。
這樣一來,即使請求失敗,程式也能顯示錯誤訊息,而不是直接中斷。
練習:觀察錯誤處理
可以將網址改成一個不存在的頁面:
url = "https://books.toscrape.com/not-found.html"
再重新執行程式。
如果網站回傳 404,raise_for_status() 就會產生 HTTP 錯誤,並由 except 區塊接住。
這個練習可以幫助我們了解,為什麼爬蟲不能只使用 requests.get(),還需要檢查回應狀態。
三、實作二:處理 HTML 欄位不存在的情況
在爬蟲中,除了 HTTP 請求失敗,另一個常見問題是找不到預期的 HTML 元素。
例如,前幾天取得書名時使用:
title = book.select_one("h3 a")["title"]
這種寫法假設 h3 a 一定存在,而且一定有 title 屬性。
如果網頁結構改變,或某筆資料缺少這個屬性,就可能出現 AttributeError 或 TypeError。
因此,可以先檢查元素是否存在,再取得資料。
from bs4 import BeautifulSoup
html = """
soup = BeautifulSoup(html, "html.parser")
book = soup.select_one("article.product_pod")
title_element = book.select_one("h3 a")
if title_element and title_element.get("title"):
title = title_element.get("title")
else:
title = None
print("書名:", title)
這裡使用 if 判斷元素是否存在,再透過 .get("title") 取得屬性。
與直接使用 ["title"] 不同,get() 在屬性不存在時會回傳 None,不會因為找不到屬性而直接產生 KeyError。
不過,取得 None 並不代表資料已經正確,因此後續仍然需要判斷這筆資料是否可以使用。
四、實作三:處理價格格式錯誤
昨天已經學會將價格中的英鎊符號移除,再透過 float() 轉換成數字。
但是,如果價格不是有效的數字,就可能出現 ValueError。
例如:
price_text = "£25.99"
price = float(price_text.replace("£", ""))
print(price)
正常情況下,這段程式碼可以將價格轉換成 25.99。
但如果取得的資料是:
price_text = "價格未知"
就無法直接轉換成浮點數。
因此,可以建立一個函式,專門處理價格的轉換。
def parse_price(price_text):
if not price_text:
return None
try:
clean_price = price_text.replace("£", "").strip()
price = float(clean_price)
if price < 0:
return None
return price
except ValueError:
return None
這個函式會先檢查價格是否為空,再嘗試移除貨幣符號並轉換成數字。
如果轉換失敗,或價格小於零,就回傳 None。
接著可以測試不同的輸入:
print(parse_price("£25.99"))
print(parse_price("£0.00"))
print(parse_price("價格未知"))
print(parse_price(None))
預期結果:
25.99
0.0
None
None
透過這個函式,就能將價格處理的邏輯集中管理,避免在每個爬蟲程式中重複撰寫相同的判斷。
五、實作四:建立資料驗證函式
完成價格處理後,接下來要建立一個函式,檢查每本書的資料是否符合預期。
這次會檢查以下項目:
書名不能為空。
價格必須是數字,而且不能小於零。
評分必須介於 1 到 5 之間。
庫存狀態必須有內容。
先建立一個資料驗證函式。
def validate_book(book):
errors = []
# 檢查書名
if not book.get("title"):
errors.append("書名缺失")
# 檢查價格
price = book.get("price")
if not isinstance(price, (int, float)):
errors.append("價格不是有效數字")
elif price < 0:
errors.append("價格不能小於零")
# 檢查評分
rating = book.get("rating")
if not isinstance(rating, int) or rating not in range(1, 6):
errors.append("評分必須介於 1 到 5")
# 檢查庫存狀態
if not book.get("availability"):
errors.append("庫存狀態缺失")
return errors
這個函式會逐一檢查每個欄位,並將發現的問題放進 errors 這個 List。
如果資料沒有問題,就會回傳空的 List。
如果資料有問題,就會回傳對應的錯誤訊息。
測試資料驗證函式
先建立一筆正常的資料。
book1 = {
"title": "Sample Book",
"price": 25.99,
"availability": "In stock",
"rating": 4
}
print(validate_book(book1))
預期結果:
[]
接著測試一筆有問題的資料。
book2 = {
"title": "",
"price": -10,
"availability": "",
"rating": 8
}
print(validate_book(book2))
這次就會得到多個錯誤訊息,因為書名和庫存狀態缺失、價格小於零,而且評分超出範圍。
這樣就能在資料進入最後的資料集之前,先檢查是否符合預期。
六、實作五:將錯誤資料與正常資料分開
前面已經完成資料驗證函式,現在可以將它整合到爬蟲流程中。
這次會建立兩個 List:
valid_books:儲存通過驗證的書籍資料。
invalid_books:儲存未通過驗證的資料及其錯誤原因。
這樣就能保留有問題的資料,而不是直接丟棄。
valid_books = []
invalid_books = []
test_books = [
{
"title": "Sample Book A",
"price": 25.99,
"availability": "In stock",
"rating": 4
},
{
"title": "",
"price": -10,
"availability": "In stock",
"rating": 8
},
{
"title": "Sample Book C",
"price": 15.50,
"availability": "In stock",
"rating": 5
}
]
for book in test_books:
errors = validate_book(book)
if errors:
invalid_books.append({
"data": book,
"errors": errors
})
else:
valid_books.append(book)
print("正常資料:", len(valid_books))
print("錯誤資料:", len(invalid_books))
這段程式碼會逐筆檢查資料。
如果 validate_book() 回傳的 List 不為空,就代表這筆資料存在問題,會被加入 invalid_books。
如果沒有錯誤,就會加入 valid_books。
這種做法可以讓爬蟲即使遇到少數格式異常的資料,也能保留其他正常資料,並且方便後續追蹤錯誤原因。
七、實作六:整合爬蟲、錯誤處理與資料驗證
前面已經分別完成請求錯誤處理、欄位檢查、價格轉換和資料驗證,現在要把這些功能整合成一個完整的爬蟲程式。
這次會從 Books to Scrape 第一頁開始,逐頁取得書籍資料,並將正常資料和錯誤資料分開儲存。
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import time
rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}
valid_books = []
invalid_books = []
visited_urls = set()
current_url = "https://books.toscrape.com/"
def validate_book(book):
errors = []
if not book.get("title"):
errors.append("書名缺失")
price = book.get("price")
if not isinstance(price, (int, float)):
errors.append("價格不是有效數字")
elif price < 0:
errors.append("價格不能小於零")
rating = book.get("rating")
if not isinstance(rating, int) or rating not in range(1, 6):
errors.append("評分必須介於 1 到 5")
if not book.get("availability"):
errors.append("庫存狀態缺失")
return errors
def parse_price(price_text):
if not price_text:
return None
try:
clean_price = price_text.replace("£", "").strip()
price = float(clean_price)
if price < 0:
return None
return price
except ValueError:
return None
while current_url:
if current_url in visited_urls:
print("發現重複網址,停止爬取")
break
visited_urls.add(current_url)
print("正在抓取:", current_url)
try:
response = requests.get(current_url, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
books = soup.select("article.product_pod")
if not books:
print("目前頁面沒有找到書籍資料")
break
for book in books:
try:
# 取得書名
title_element = book.select_one("h3 a")
title = (
title_element.get("title")
if title_element
else None
)
# 取得價格
price_element = book.select_one(".price_color")
price_text = (
price_element.get_text(strip=True)
if price_element
else None
)
price = parse_price(price_text)
# 取得庫存狀態
availability_element = book.select_one(
".availability"
)
availability = (
availability_element.get_text(strip=True)
if availability_element
else None
)
# 取得評分
rating_element = book.select_one(".star-rating")
rating = None
if rating_element:
classes = rating_element.get("class", [])
for class_name in classes:
if class_name in rating_map:
rating = rating_map[class_name]
break
# 建立書籍資料
book_info = {
"title": title,
"price": price,
"availability": availability,
"rating": rating
}
# 驗證資料
errors = validate_book(book_info)
if errors:
invalid_books.append({
"data": book_info,
"errors": errors
})
else:
valid_books.append(book_info)
except (AttributeError, TypeError, KeyError) as e:
invalid_books.append({
"data": {},
"errors": [str(e)]
})
# 尋找下一頁
next_link = soup.select_one("li.next a")
if next_link:
next_href = next_link.get("href")
current_url = urljoin(current_url, next_href)
else:
current_url = None
time.sleep(1)
except requests.exceptions.RequestException as e:
print("請求失敗:", e)
break
print("爬取完成")
print("正常資料:", len(valid_books))
print("錯誤資料:", len(invalid_books))
程式碼解說
這段程式碼將前面學過的功能整合在一起。
首先,使用 Requests 取得網頁,再透過 BeautifulSoup 找出書籍區塊。
接著,逐筆擷取書名、價格、庫存狀態和評分,並將資料整理成 Dictionary。
每筆資料都會經過 validate_book() 檢查。如果資料符合條件,就會加入 valid_books;如果不符合,就會加入 invalid_books,並保留錯誤原因。
最後,程式會尋找 Next 按鈕,取得下一頁網址,繼續爬取直到沒有下一頁為止。
這樣就能在同一個爬蟲流程中完成資料擷取、錯誤處理和資料驗證。
八、實作七:查看錯誤資料與驗證結果
完成爬蟲後,可以先確認正常資料與錯誤資料的數量。
print("正常資料筆數:", len(valid_books))
print("錯誤資料筆數:", len(invalid_books))
接著查看錯誤資料的內容。
for item in invalid_books[:10]:
print("資料:", item["data"])
print("錯誤原因:", item["errors"])
print("-" * 40)
這樣就能查看前 10 筆未通過驗證的資料,以及每筆資料的錯誤原因。
如果錯誤資料數量很多,就可以進一步檢查是不是網站的 HTML 結構改變,或是某個欄位的擷取方式有問題。
如果沒有錯誤資料,也代表目前取得的資料都通過了這次設定的驗證條件,但仍不代表所有資料都一定正確。
九、實作八:將正常資料與錯誤資料分別儲存成 CSV
完成資料驗證後,最後要將正常資料和錯誤資料分開保存。
with open(
"valid_books.csv",
"w",
newline="",
encoding="utf-8-sig"
) as file:
fieldnames = [
"title",
"price",
"availability",
"rating"
]
writer = csv.DictWriter(
file,
fieldnames=fieldnames
)
writer.writeheader()
writer.writerows(valid_books)
print("正常資料已儲存")
這個檔案只會包含通過驗證的書籍資料,方便後續使用 Pandas 進行分析。
儲存錯誤資料
with open(
"invalid_books.csv",
"w",
newline="",
encoding="utf-8-sig"
) as file:
fieldnames = ["data", "errors"]
writer = csv.DictWriter(
file,
fieldnames=fieldnames
)
writer.writeheader()
writer.writerows(invalid_books)
print("錯誤資料已儲存")
這個檔案會保留未通過驗證的資料和錯誤原因。
將兩種資料分開保存,可以避免錯誤資料直接進入後續分析,也能保留需要進一步檢查的內容。